Ocala
M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding
Cho, Jaemin, Mahata, Debanjan, Irsoy, Ozan, He, Yujie, Bansal, Mohit
Document visual question answering (DocVQA) pipelines that answer questions from documents have broad applications. Existing methods focus on handling single-page documents with multi-modal language models (MLMs), or rely on text-based retrieval-augmented generation (RAG) that uses text extraction tools such as optical character recognition (OCR). However, there are difficulties in applying these methods in real-world scenarios: (a) questions often require information across different pages or documents, where MLMs cannot handle many long documents; (b) documents often have important information in visual elements such as figures, but text extraction tools ignore them. We introduce M3DocRAG, a novel multi-modal RAG framework that flexibly accommodates various document contexts (closed-domain and open-domain), question hops (single-hop and multi-hop), and evidence modalities (text, chart, figure, etc.). M3DocRAG finds relevant documents and answers questions using a multi-modal retriever and an MLM, so that it can efficiently handle single or many documents while preserving visual information. Since previous DocVQA datasets ask questions in the context of a specific document, we also present M3DocVQA, a new benchmark for evaluating open-domain DocVQA over 3,000+ PDF documents with 40,000+ pages. In three benchmarks (M3DocVQA/MMLongBench-Doc/MP-DocVQA), empirical results show that M3DocRAG with ColPali and Qwen2-VL 7B achieves superior performance than many strong baselines, including state-of-the-art performance in MP-DocVQA. We provide comprehensive analyses of different indexing, MLMs, and retrieval models. Lastly, we qualitatively show that M3DocRAG can successfully handle various scenarios, such as when relevant information exists across multiple pages and when answer evidence only exists in images.
MiMiC: Minimally Modified Counterfactuals in the Representation Space
Singh, Shashwat, Ravfogel, Shauli, Herzig, Jonathan, Aharoni, Roee, Cotterell, Ryan, Kumaraguru, Ponnurangam
Language models often exhibit undesirable behaviors, such as gender bias or toxic language. Interventions in the representation space were shown effective in mitigating such issues by altering the LM behavior. We first show that two prominent intervention techniques, Linear Erasure and Steering Vectors, do not enable a high degree of control and are limited in expressivity. We then propose a novel intervention methodology for generating expressive counterfactuals in the representation space, aiming to make representations of a source class (e.g., "toxic") resemble those of a target class (e.g., "non-toxic"). This approach, generalizing previous linear intervention techniques, utilizes a closed-form solution for the Earth Mover's problem under Gaussian assumptions and provides theoretical guarantees on the representation space's geometric organization. We further build on this technique and derive a nonlinear intervention that enables controlled generation. We demonstrate the effectiveness of the proposed approaches in mitigating bias in multiclass classification and in reducing the generation of toxic language, outperforming strong baselines.
Does Outrage Signal Cyber Attacks? Predicting "Bad Behavior" from Sentiment in Online Content
Hollingshead, Kristy (Florida Institute for Human and Machine Cognition) | Dorr, Bonnie J. (Florida Institute for Human and Machine Cognition) | Dalton, Adam (Florida Institute for Human and Machine Cognition) | Barton, Meg (Leidos, Inc.)
We demonstrate that it is possible to leverage big data in the form of tweets and linked webpages to find expressions of sentiment that signal "bad behavior" such as cyber attacks. We hypothesize that expressions of "outrage" (high intensity, negative affect sentiment) against an organization in public data may be predictive of cyber attacks for two reasons: 1) threat actors may be motivated to launch an attack based on anger/discontent, and 2) outrage associated with an organization or industry may increase the likelihood of that organization or industry being victimized by threat actors (i.e., as a form of "vigilante justice"). We measure sentiment in online content and determine trends in public emotion and their correlation to trends in cyber attacks, as reported in Hackmageddon. We demonstrate that dimensions of sentiment, as afforded by our use of the Circumplex model of emotion, do yield correlations to reported cyber attacks, but differ dependent upon the domain of the data. Thus the use of this technique requires careful analysis for optimal application.
Cyberdyne's HAL Exoskeleton Helps Patients Walk Again in First Treatments at U.S. Facility
Danny Bal was riding his brand new motorcycle to work from his home in Ocala, Florida two years ago when the driver of an oncoming car fell asleep and ploughed into Bal's electric-blue bike. After the accident, which crushed three of Bal's thoracic vertebrae and shredded a spinal nerve, Bal adjusted to life in a wheelchair. He added a motorized lift to his beloved F-250 truck, explored local trails with a hand-powered bike, and joined a therapeutic horseback riding program. Now, one of Bal's daughters is about to get married, and 57-year-old Bal wants to walk in her ceremony. So on a recent Friday morning in December at Brooks Rehabilitation in Jacksonville, Florida, Bal was back on his feet, taking slow but steady steps as his granddaughter cheered from the sidelines.
An Ostrich-Like Robot Pushes the Limits of Legged Locomotion
What looks like a tiny mechanical ostrich chasing after a car is actually a significant leap forward for robot-kind. The clever and simple two-legged robot, known as the Planar Elliptical Runner, was developed at the Institute for Human and Machine Cognition in Ocala, Florida, to explore how mechanical design can be used to enable sophisticated legged locomotion. A video produced by the researchers shows the robot being tested in a number of situations, including on a treadmill and running behind and alongside a car with a helping hand from an engineer. In contrast to many other legged robots, this one doesn't use sensors and a computer to help balance itself. Instead, its mechanical design provides dynamic stability as it runs.
Government regulators are looking into fatal Tesla crash involving Autopilot
Tesla announced today that the National Highway Traffic Safety Administration has opened an investigation into a recent fatal crash of a Model S with the company's Autopilot feature activated. The accident took place on May 7th in a small West Florida town called Williston. The Florida Highway Patrol is also conducting its own investigation of the accident, according to a public affairs officer there. The same officer reported that Tesla has, since the fatal accident in May, sent engineers down to Ocala, Florida to assist investigators in accessing data they needed to evaluate the causes of the crash. Tesla offered an account of the event in a blog post titled "A Tragic Loss" that went up today, detailing the crash, an "extremely rare circumstance," which occurred on a divided highway.